PCA as a practical indicator of OPLS-DA model reliability.

نویسندگان

  • Bradley Worley
  • Robert Powers
چکیده

BACKGROUND Principal Component Analysis (PCA) and Orthogonal Projections to Latent Structures Discriminant Analysis (OPLS-DA) are powerful statistical modeling tools that provide insights into separations between experimental groups based on high-dimensional spectral measurements from NMR, MS or other analytical instrumentation. However, when used without validation, these tools may lead investigators to statistically unreliable conclusions. This danger is especially real for Partial Least Squares (PLS) and OPLS, which aggressively force separations between experimental groups. As a result, OPLS-DA is often used as an alternative method when PCA fails to expose group separation, but this practice is highly dangerous. Without rigorous validation, OPLS-DA can easily yield statistically unreliable group separation. METHODS A Monte Carlo analysis of PCA group separations and OPLS-DA cross-validation metrics was performed on NMR datasets with statistically significant separations in scores-space. A linearly increasing amount of Gaussian noise was added to each data matrix followed by the construction and validation of PCA and OPLS-DA models. RESULTS With increasing added noise, the PCA scores-space distance between groups rapidly decreased and the OPLS-DA cross-validation statistics simultaneously deteriorated. A decrease in correlation between the estimated loadings (added noise) and the true (original) loadings was also observed. While the validity of the OPLS-DA model diminished with increasing added noise, the group separation in scores-space remained basically unaffected. CONCLUSION Supported by the results of Monte Carlo analyses of PCA group separations and OPLS-DA cross-validation metrics, we provide practical guidelines and cross-validatory recommendations for reliable inference from PCA and OPLS-DA models.

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

ropls: PCA, PLS(-DA) and OPLS(-DA) for multivariate analysis and feature selection of omics data

4 Hands-on 3 4.1 Loading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 4.2 Principal Component Analysis (PCA) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 4.3 Partial least-squares: PLS and PLS-DA . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 4.4 Orthogonal partial least square...

متن کامل

1,5-Anhydroglucitol, an indicator of short term glycaemic control, is the most discriminatory metabolomic marker in adolescents with type 1 diabetes compared to control subjects

Methods Design Case control study. Setting Tertiary paediatric hospital clinic. Population 27 (14F/13M) adolescents with T1D (age (median, interquartile range) 15.5, 14.7-16.4 years; duration 7.7; 6.0-11.8 years; HbA1c 9.1, 8.1-10.1%); glucose 13.35 (7.60-17.85) and 27 (14F/13M) control participants (age 15.1, 14.4-16.8 years). BMI was <95 percentile. Measures Fasting plasma and urine metabolom...

متن کامل

Efficient Discovery of Quality Control Markers for Gastrodia elata Tuber by Fingerprint-Efficacy Relationship Modelling.

INTRODUCTION Gastrodia elata tuber (GET) has been widely used in China as a famous herbal medicine. However, the quality control markers (QCMs) for GET still need further investigation. OBJECTIVE To develop a rational strategy based on fingerprint-efficacy relationship modelling to discover the efficacy-related QCMs, using GET as a case study. METHODOLOGY The high-performance liquid chromat...

متن کامل

1H NMR-Based Metabonomic Study of Functional Dyspepsia in Stressed Rats Treated with Chinese Medicine Weikangning

1H NMR-based metabolic profiling combined with multivariate data analysis was used to explore the metabolic phenotype of functional dyspepsia (FD) in stressed rats and evaluate the intervention effects of the Chinese medicine Weikangning (WKN). After a 7-day period of model establishment, a 14-day drug administration schedule was conducted in a WKN-treated group of rats, with the model and norm...

متن کامل

A Pilot Metabolic Profiling Study of Patients With Neonatal Jaundice and Response to Phototherapy

Phototherapy has been widely used in treating neonatal jaundice, but detailed metabonomic profiles of neonatal jaundice patients and response to phototherapy have not been characterized. Our aim was to depict the serum metabolic characteristics of neonatal jaundice patients relative to controls and changes in response to phototherapy. A (1) H nuclear magnetic resonance (NMR)-based metabonomic a...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

عنوان ژورنال:
  • Current Metabolomics

دوره 4 2  شماره 

صفحات  -

تاریخ انتشار 2016